Papers with French language

17 papers
French GossipPrompts: Dataset For Prevention of Generating French Gossip Stories By LLMs (2024.eacl-short)

Copied to clipboard

Challenge: Large Language Models (LLMs) are undergoing a dynamic transformation . however, there is a potential risk of LLMs creating gossips when prompted with contexts .
Approach: a dataset is used to identify prompts that lead to the creation of gossipy content in the french language.
Outcome: a new dataset identifies prompts that lead to the creation of gossipy content in the french language . the model achieves an accuracy of 89.95% .
Multi-lingual neural title generation for e-Commerce browse pages (N18-3)

Copied to clipboard

Challenge: e-Commerce websites are automatically generating millions of browse pages . manual creation of titles is infeasible due to the huge number of browse page types .
Approach: They propose to use sequence-to-sequence models to generate titles for languages . they train the models on multi-lingual data, thereby creating one joint model .
Outcome: The proposed model can generate titles in three different languages, with a focus on low-resource French.
Providing Semantic Knowledge to a Set of Pictograms for People with Disabilities: a Set of Links between WordNet and Arasaac: Arasaac-WN (2020.lrec-1)

Copied to clipboard

Challenge: Pictograms are a tool that is increasingly used by people with cognitive or communication disabilities.
Approach: They propose a database that links WordNet and Arasaac to link pictograms to semantic knowledge.
Outcome: The proposed database links pictograms with WordNet and Arasaac to create language-independent prototypes.
A Computational Analysis and Exploration of Linguistic Borrowings in French Rap Lyrics (2024.acl-srw)

Copied to clipboard

Challenge: rap is a popular genre in the u.s. and has been used in countries far beyond the uk . linguistic borrowings are especially intriguing in countries such as the eu and europe .
Approach: They manually annotate a lexicon of over 700 borrowings in the French language . they find that there are increases in the proportion of linguistic borrowings, interjections, and Niger-Congo borrowings .
Outcome: The proposed method analyzes a corpus of over 8000 french rap song lyrics and shows that rap borrowings are increasing in prevalence and interjections are decreasing.
CoFiF Plus: A French Financial Narrative Summarisation Corpus (2022.lrec-1)

Copied to clipboard

Challenge: Existing corpora for financial narrative summarisation exists in English but there is a significant lack of financial text resources in the French language.
Approach: They propose to use natural language processing to analyse financial documents to find the best summarisation methods.
Outcome: The proposed dataset is the first to provide a comprehensive set of financial text written in French.
Speech Resources in the Tamasheq Language (2022.lrec-1)

Copied to clipboard

Challenge: In this paper, we present two datasets for Tamasheq, a developing language mainly spoken in Mali and Niger . we share unlabeled audio data in five languages: french, Fulfulde, Hausa, Tamaheq and Zarma .
Approach: They present two datasets for Tamasheq, a developing language mainly spoken in Mali and Niger.
Outcome: The proposed datasets are used in the IWSLT 2022 low-resource speech translation track . they consist of radio recordings from daily broadcast news in Niger and Mali .
FQuAD2.0: French Question Answering and Learning When You Don’t Know (2022.lrec-1)

Copied to clipboard

Challenge: Question Answering, including Reading Comprehension, has seen significant scientific breakthroughs over the past few years . but most of these breakthroughs are centered on the English language .
Approach: They propose a dataset to train Question Answering models in the French language . they extend the dataset to 17,000+ unanswerable questions annotated adversarially .
Outcome: The proposed dataset makes it possible to train French Question Answering models with the ability to distinguish unanswerable questions from answerable ones.
NLP Analytics in Finance with DoRe: A French 250M Tokens Corpus of Corporate Annual Reports (2020.lrec-1)

Copied to clipboard

Challenge: Recent advances in neural computing and word embeddings for semantic processing open many new applications areas which had been left unaddressed due to inadequate language understanding capacity.
Approach: They propose a French and dialectal French corpus for NLP analytics in finance, regulation and investment.
Outcome: The proposed corpus is designed to be as modular as possible to allow for maximum reuse in different tasks pertaining to Economics, Finance and Investment.
Know When to Fuse: Investigating Non-English Hybrid Retrieval in the Legal Domain (2025.coling-main)

Copied to clipboard

Challenge: Existing research focuses on a limited set of retrieval methods, evaluated in pairs on domain-general datasets exclusively in English.
Approach: They evaluate the efficacy of hybrid search across a variety of retrieval models in the french language . they find that fusion of different domain-general models consistently enhances performance .
Outcome: The proposed model improves in-domain performance compared to a single model in a zero-shot context . the proposed model also improves when the models are trained in- domain .
A Benchmark of French ASR Systems Based on Error Severity (2025.coling-main)

Copied to clipboard

Challenge: Automatic Speech Recognition (ASR) transcription errors are often assessed using metrics that compare them with a reference transcription.
Approach: They propose to categorize transcription errors into four levels of severity based on objective linguistic criteria, contextual patterns, and the use of content words as the unit of analysis.
Outcome: The proposed evaluation categorizes errors into four levels of severity based on objective linguistic criteria, contextual patterns, and the use of content words as the unit of analysis.
GeNRe: A French Gender-Neutral Rewriting System Using Collective Nouns (2025.findings-acl)

Copied to clipboard

Challenge: Gender rewriting is an NLP task that uses gendered forms to mitigate gender biases.
Approach: They propose a French gender-neutral rewriting system using collective nouns, which are gender-fixed in French.
Outcome: The proposed system detects gendered forms and replaces them with neutral or opposite forms.
PAGnol: An Extra-Large French Generative Model (2022.lrec-1)

Copied to clipboard

Challenge: a growing number of pre-trained language models are available in many different languages.
Approach: They propose a French-language GPT model with scaling laws to train it efficiently . they evaluate the models on discriminative and generative tasks in French .
Outcome: The proposed model trains with the same computational budget as CamemBERT, a model 13 times smaller.
Introducing RezoJDM16k: a French KnowledgeGraph DataSet for Link Prediction (2022.lrec-1)

Copied to clipboard

Challenge: Knowledge graphs are used for information extraction, search engines, question answering, and recommendation systems.
Approach: They propose a French knowledge graph dataset based on RezoJDM.
Outcome: The proposed dataset can be used in many downstream tasks for the French language . it shows that it embeds knowledge graph baselines for link prediction tasks .
FRACAS: a FRench Annotated Corpus of Attribution relations in newS (2024.lrec-main)

Copied to clipboard

Challenge: Quotation extraction is a useful task, but it is not widely studied in other languages.
Approach: They propose to annotate a manually annotated corpus of 1,676 newswire texts in French for quotation extraction and source attribution.
Outcome: The proposed system is compared to the most recent system for quotation extraction in the French language.
FReND: A French Resource of Negation Data (2024.lrec-main)

Copied to clipboard

Challenge: Negation data are limited by the language models of the BERT-generation, which are still underperforming on tasks and benchmarks featuring negation.
Approach: FReND is a freely available corpus of French language in which negations are hand-annotated by their cues and scopes.
Outcome: FReND is the largest dataset available for french negation analysis . it is a valuable resource for linguistic research and as training data for AI tasks such as negation detection.
French Tweet Corpus for Automatic Stance Detection (2020.lrec-1)

Copied to clipboard

Challenge: a new corpus of tweets is being developed for automatic stance detection of fake news . the task involves determining the attitude expressed in a text toward a target . this is a difficult task to overcome as discussions about fake news are controversial .
Approach: They propose to build a human-annotated corpus for automatic stance detection of tweets in french . they propose to use four classes broadly adopted by the community for annotation .
Outcome: The proposed corpus is the first freely available stance annotated tweet corpus in the french language.
DrBERT: A Robust Pre-trained Model in French for Biomedical and Clinical domains (2023.acl-long)

Copied to clipboard

Challenge: Recent studies have shown that pre-trained language models improve performance on a wide range of NLP tasks.
Approach: They propose to use pre-trained language models to train medical domains on French language to compare performance with specialized ones.
Outcome: The proposed models can take advantage of existing biomedical models in a foreign language by further pre-training them on our targeted data.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations